Skip to content

feat(diagrams): the agent draws in the chat — render_diagram, phase 1 of the Rich Diagrams spec - #104

Merged
ndemianc merged 15 commits into
developfrom
feat/rich-diagrams
Oct 8, 2026
Merged

ndemianc merged 15 commits into
developfrom
feat/rich-diagrams

Conversation

@ndemianc

@ndemianc ndemianc commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

What this is

Ask the agent how something is put together and it answers with boxes and arrows typed out of characters. They break with the font, the theme and the width of the panel; they cannot be clicked or exported.

With this PR the agent has a tool, render_diagram. The model describes structure — nodes, edges, groups, one accent, never a coordinate or a colour. The editor validates the description, lays it out in one house style and paints it in the chat, in the editor's theme, with nodes that open the code they stand for.

This is phase 1 of the Rich Diagrams spec: Graph JSON only. Mermaid, Vega-Lite and raw SVG are not built.

The architecture fixture, light palette

The same fixture, dark palette

Two of the committed snapshots: what the SVG export writes. In the chat the colours are the editor's own theme tokens.

How it works

  1. A request from a client that can render carries the tool and a short block of rules.
  2. The model calls the tool. A placeholder holds the place, titled as soon as the title has streamed.
  3. The host makes the lossless fixes and validates. Valid: a record is stored and sent to the chat. Invalid: every error goes back to the model, once.
  4. The chat validates the record again, measures the text in the real font, lays it out and paints it from allow-lists.
  5. If the one repair pass does not produce a valid spec, the editor draws what is valid and says what it left out — or shows the errors, the source and Retry. Never an empty card.
Where What
diagram/schema.js validate.js repair.js The versioned schema, every error at once with a JSON Pointer, and the repair ladder
diagram/theme.js layout.js The style guide as numbers; a layered layout with orthogonal connectors
diagram/scene.js text.js ascii.js The painter (two allow-lists, textContent), the screen-reader outline, Mermaid and source export, the text fallback
diagram/tool.js service.js The tool and its rules; one conversation's diagrams and their one repair pass
diagram/links.js exportCheck.js stats.js bundle.js Links resolved inside the workspace; exports checked before they are saved; local counters; the shared modules inlined into the page
agent.js, providers/* The tool in the loop; a call's id and its streamed arguments
sessionEvents.js sessions.js extension.js Diagrams stored with the turn and replayed; stubs at compaction and get_diagram; link, export, retry
media/chat.html The card, its states, zoom, toolbar, fallback

The page's content security policy is unchanged. The shared modules are inlined under the nonce the page already has, and it still loads nothing.

Where it departs from the spec

The spec This PR Why
Graph JSON, Mermaid, Vega-Lite, opt-in raw SVG Graph JSON only Phase 1. The rules tell the model to use a table for numbers and a numbered list for sequences, not a format the chat would show as source
Layout by ELK.js An in-house layered layout ELK is EPL-2.0 in an otherwise MIT-clean repository, about 1.5 MB, and the extension has no build step. layout.layout() is the one entry point, so ELK can replace it
Validation by Ajv A small interpreter for the subset of JSON Schema the schema uses The same reasons
"Dedupe ids" and "drop edges to unknown nodes" are fixes made without a model Lossy fixes run only after the model's one repair pass has failed The spec's own table says semantic errors go to the model, "because intent is needed". An edge to billing when the node is bill is a typo; dropping it silently changes what the diagram says
The host accepts two actions from a diagram Four: openLink, export, retry, ascii The spec's own Retry button and text fallback need the other two. Each names a diagram by id; none carries a path or a URL
Rendering in a worker or sandboxed frame In the chat webview For Graph JSON the painter executes nothing the model wrote. Mermaid and Vega run third-party parsers over model text and should get the frame

The rule the ladder follows: an auto-fixed diagram never says something different from what the model wrote; a degraded one always says what it lost.

Also in here: a chat rendering bug, unrelated to diagrams

The first two commits fix a bug that is on develop today. An answer that mentioned a code fence in passing — the model quoted three backticks inside an inline span of four — was rendered as two one-character code blocks, and the rest of the paragraph came out one streamed fragment per line.

The renderer treated any three backticks, anywhere, as a fence. And once a "closing" fence had no newline after it, the streaming renderer froze every delta into its own block. A fence is now a line of its own, inline code can be delimited by any number of backticks, and only a block whose closing line is complete is frozen.

96399b6 pins what the renderer did before; 3b05fd3 is the fix. Replayed in Chromium with the session's exact text in 184 deltas: the paragraph was 57 lines and two code blocks before, and is one paragraph and none after. The two commits stand alone, and can be cherry-picked into their own PR if they should merge first.

The commits

Commit What
96399b6 Pins how the chat renders fences and inline code
3b05fd3 The chat fix
42159e5 The diagram schema, validator and repair ladder
fb2e67a The theme and the layout
f35fbc1 The painter, the text outline, the text fallback
ce02534 render_diagram in the agent, the host and the chat
822cfc2 The browser check, the real-editor check, the eval harness
db68f68 docs/RICH-DIAGRAMS.md, CLAUDE.md, the fixtures' README
a317a45 Review: the width estimate for an emoji, and a dead store
e84b368 Review: names every object answers to, as a model id, a schema version or a shape
49d16b8 Review: which call repairs which diagram, replaced drawings, the order of a replay
51e4490 Review: the full-size view is modal, the picture is not a button, one diagram is one card
0f55187 A saved SVG carries no links
3c51f11 The doc after the review round

Each one passes ./scripts/test-extensions.sh on its own. The gate was run on every committed tree in a separate checkout, not only on the tip.

Review

Nine findings in two rounds: two from the code-quality check and seven from Copilot. All nine were right. Each is fixed with a test that fails on the code as reviewed, and each thread is answered and resolved.

Finding Outcome
theme.js: c > 0xffff is always false The width estimate asked "full-width script" before "emoji", so an emoji was measured at 1.0 em and not 1.1. Asked the other way round (a317a45)
layout.js: a value assigned and never read Removed (a317a45)
stats.js: a model called constructor or __proto__ The lookup found Object, or the prototype every object shares, and wrote a counter to it. Keys are read and written as own entries (e84b368)
validate.js: a schema version called toString validate, prepare and accept threw where they promise an error. A version must be one of the numbers this editor reads, which also stops the string "1" passing as version 1 (e84b368)
agent.js: a repair written as loose JSON loses its identity The first attempt was drawn degraded and the correction as a second card. A call's identity is read from the spec as the ladder normalizes it, whatever form it arrived in (49d16b8)
service.js: stubs for drawings that were replaced A replaced drawing gets no stub and is not a known id; get_diagram on it names the drawing the chat shows (49d16b8)
sessionEvents.js: prose collapsed ahead of the diagram A replay and the Markdown export keep prose, picture, prose (49d16b8)
chat.html: the full-size view is modal in name only The page behind it is inert, Tab stays inside, a linked node opens on Enter, and closing returns the focus (51e4490)
chat.html: the picture is a button with links inside it The picture is not a control. "Full size" is a button in the toolbar; a click on the picture still works for the pointer (51e4490)

Fixing these turned up six faults nobody had reported. They are fixed in the same commits, and the last has one of its own.

  • A redraw left both drawings on screen, where a reopened session shows one.
  • A different diagram arriving mid-repair painted over the one that was waiting. Two diagrams on file, one on screen.
  • A tool call named __proto__ made the page throw, and a later diagram was never shown.
  • "shape": "constructor" found Object in the synonym table. The model was told "expected a string, got function" about a string it had written.
  • The Markdown export put a diagram with no prose before it under "You".
  • A saved SVG kept its nodes' link roles and tab stops, which do nothing in a file (0f55187).

The first three are in the chat page, and neither net could have caught them. The gate only pinned the page's source, and the browser check never played messages in the order a real run sends them: a placeholder at the start of every call, before anyone knows whether it is a new diagram, a repair or a redraw. Both gaps are closed. The page's card bookkeeping now runs in the gate, and the browser check plays seven runs in the real order, then replays each as a reopened session and compares the two.

Verification

The numbers are for the tip, 3c51f11.

  • The gate: 62 suite files, on macOS arm64 with Node 24, and on Linux (Ubuntu 22.04 arm64, Node 18) in a container with no network and a read-only checkout. Merged with develop at edc2752 in a scratch checkout, 63 pass.
  • CI: Extension unit tests and the CodeQL analysis pass on GitHub's runner for 3c51f11.
  • Twelve diagram suites, 270 tests. They include a 1,500-spec layout fuzz, 39 seeded broken specs with the exact outcome of each attempt, SVG snapshots in both palettes, the gallery at chat-panel widths, the real runAgent loop against a scripted provider, extension.js's own functions sliced out and run against a stand-in for vscode, and the page's card bookkeeping run against a stand-in for the few DOM calls it makes.
  • The chat page in a browser (scripts/diagram-browser-check.js): the shipped chat.html in headless Chrome under the real policy, 186 checks. Hostile labels are on the page as text and nothing ran; no policy violation; no request. The full-size view's focus is checked in real Chrome, and seven runs are played in the real order, live and then reopened.
  • The real editor (scripts/diagram-editor-check.js): a second, throwaway instance of the dev build loads this branch's extension and talks to a stand-in provider on localhost, 19 checks. The tool and its rules reach the model, the diagram is painted in the real webview in the editor's colours, the full-size view opens from its button and is modal there, a linked node opens its file at the symbol, and the session on disk holds the diagram.
  • Mutation testing, before the review: 168 single-edit defects seeded across the modules, the host glue, the page, the eval harness and the chat fix. Twelve survived at first: nine were gaps in the suites, now closed, and three were redundant code, now removed. All 165 that still applied were caught.
  • Mutation testing, the review round: 69 more against what the round changed, 57 against the gate and 12 against the browser check. Two survived at first, both gaps in the card tests, now closed. All 69 are caught. The first set was not run again after the round.
  • Timing: the diagram suites and the chat suite also pass with every timer delayed by 15 ms.

Not verified

  • No live model has drawn a diagram. Every "model" above is a script. scripts/diagram-eval.js is the spec's eval — 30 prompts that should produce a diagram and 10 that should not, run through the real agent loop — and it has not been run, because --run makes billed calls. Nothing here says how often a real model's first attempt is valid, or whether it reaches for the tool at all.
  • A screen reader. The accessibility fixes were checked through focus and inert, in Chrome and in the editor's own webview. Nobody listened to one.

For the reviewer to decide

  • On by default. levelcode.ai.diagrams.enabled defaults to true, and costs about 970 tokens of tool and rules on every agent request while it is on. It is constant for a session, so it sits in the cached prefix. Flip the default if this should ship dark until the eval has run.
  • A model can be switched off with diagrams: false in its providers/catalog.js row. That is the lever for one that fails the eval.
  • docs/RICH-DIAGRAMS.md is an implementation record, not the spec. The spec is a private document.

Known limits

Measured over 6,000 random specs at the model-facing limits. None of the nine gallery diagrams shows any of the three, at its natural width or at 560, 420 and 320 px.

Limit How often
An edge label falls back to a halo over a line 6.5% of specs, mostly dense fan-ins in a top-to-bottom flow
Two connectors swap lanes in one gap 0.3%
A group's name has a connector passing behind it 2.7% of specs with groups

Each is bounded in the layout suite, so it cannot get worse quietly. The rest are in the doc.

One more, from the review round: while a run is live, a repaired diagram is drawn where its first attempt was waiting, and a reopened session shows it at the call that got it right. The cards are the same; the place differs by whatever the model wrote in between.

To try it

On this branch, in the checkout that has vscode/, ./scripts/run-dev.sh is enough. From a git worktree, which has no vscode/ of its own, load the worktree's extension into the build that exists:

./scripts/run-dev.sh --extensionDevelopmentPath=/absolute/path/to/worktree/extensions/levelcode-ai

The title bar says [Extension Development Host]. In agent mode, ask something whose answer is structure: "how does a request flow through this app?"

…fore changing it

The chat's Markdown renderer lives inline in media/chat.html, and no suite ran
it. The next commit changes how it finds a code fence, so this one writes down
what it does today, against today's code:

  an ordinary fenced block      a code block; the language line is not in it
  two blocks, prose between     and a block still being typed is shown as code
  inline code                   escaped, never formatted; a lone backtick is text
  a path in inline code         a file chip, when it names a real project file
  a fence inside a list item    still a code block, indented as it was
  a fence stuck to a line       "Run this: ```bash" opens a block, and a closing
                                fence stuck to the last line of code closes one
                                — sloppy, and already tolerated

The functions are sliced out of chat.html itself, the way shHighlight.test.js
does it, so the suite runs the shipped code and not a copy of it.
…aragraph came out a fragment per line

An answer that mentioned a code fence in passing — "each closed by a plain
```` ``` ````" — was rendered as two code blocks holding one backtick each, and
the rest of the paragraph arrived one streamed fragment per line: "gr",
"ounded ent", "irely in the actual code". Reported from a real session. The
model had quoted the fence correctly, in an inline span of four backticks.

Two things were wrong, and the second is what made it look so bad.

render() split the text on every run of three backticks, wherever it stood. So
```` ``` ```` was three fences: a block containing "` ", and then an unfinished
block containing the rest of the message.

lastStableIndex() decides what the streaming renderer may freeze — move out of
the live tail and into the page for good. It counted fences the same way, and
when the text after an even-numbered one had no newline yet, it answered "all
of it". From then on every delta was frozen as it arrived, each in its own
<div>.

A fence is a line. mdSegments() reads the text line by line: a block opens on a
line that is, after its indentation, three or more backticks and a language (or
anything else without a backtick in it), and closes on a line ending in a run
at least as long as the one that opened it. render() and lastStableIndex() both
use it, so they cannot disagree again. A block counts as finished only once its
closing line has its newline; text after it stays live until the next block
finishes.

Inline code takes the count seriously too: a run of N backticks opens a span
that the next run of exactly N closes, on the same line. That is how three
backticks are quoted (inside four), and one (inside two).

Kept, because models do it and the old code let them: a fence stuck to the end
of a line of prose ("Run this: ```bash") still opens a block — when nothing but
a one-word language follows it and it is not the closing half of an inline
span — and a closing fence stuck to the last line of code still closes one.

Not changed: tildes are not fences; and a line of prose that ENDS in three bare
backticks still opens a block, as it did. CommonMark would not; a model that
wants to say "```" has an inline span for it, and now it works.

chatMarkdownFences.test.js gains the report itself, whole and streamed — in
sixty different chunkings, and one character at a time; the property that
makes streaming safe (streamed and whole are the same page, and only finished
blocks are frozen) over six documents and 400 random mixes of backticks,
newlines and words; N-backtick spans; fences nested by length; language lines
that say more than a word; and that nothing in a message becomes markup. The
six pins from the last commit pass unchanged. The ten new tests fail on the old
code.

Replayed in Chromium against the shipped page with the session's exact text, in
184 deltas: the paragraph was 57 lines and two code blocks before, and is one
paragraph and none after. 21 single-edit mutations of the new code are each
caught by the suite.
…rror at once, and the repair ladder

First slice of rich diagrams; docs/RICH-DIAGRAMS.md arrives with the last one.
The agent draws flows and architectures out of box-drawing characters. The plan
is that it describes STRUCTURE instead — nodes, edges, groups, one accent,
never a coordinate or a colour — and the editor draws. This commit is the part
that decides whether a description may be drawn. Nothing calls it yet.

  diagram/schema.js     the versioned schema, and a small interpreter for the
                        subset of JSON Schema it uses. One object is both the
                        tool's input schema and what a spec is checked against.
                        Two tiers of limits: HOUSE is what a model is held to
                        (12 nodes; labels 28, second lines 32, edge labels 20,
                        title 80; groups two deep), HARD is what the renderer
                        will still draw (24 nodes, 48 edges).
  diagram/validate.js   schema, then semantics: edge ends exist, ids are unique,
                        one accent, groups known, acyclic, at most two deep.
                        EVERY error, each with a JSON Pointer, what was
                        expected and the valid options:
                          /edges/1/to: unknown node "billing". Known ids: in, jev, bill, rev.
  diagram/repair.js     the ladder. prepare() takes one call up it.

The ladder departs from the spec in one place, on purpose. The spec lists
"dedupe ids" and "drop edges to unknown nodes" among the fixes made without a
model, while its own table says semantic errors go to the model, "because
intent is needed". Both cannot hold: an edge to "billing" when the node is
"bill" is a typo the model fixes in one pass, and dropping it silently changes
what the diagram says. So:

  1  auto-fix, lossless only   lenient JSON (comments, trailing commas, smart
                               quotes, bare keys), synonyms ("diamond" for
                               decision, source/target for from/to), slugged
                               ids, an edge that names a node by its label,
                               long labels shortened with the full text kept
  2  every error, once         returned for the model's one repair pass
  3  degrade                   only now the lossy fixes: edges to nowhere
                               dropped, duplicates renamed, one accent kept,
                               counts up to the HARD tier — each loss named

An auto-fixed diagram never says something different from what the model
wrote; a degraded one always says what it lost.

accept() is the same gate for a spec that is read back rather than received —
from a session file, say: validate, hand on only the fields the schema
declares, and fall back to the rungs that need no model.

No Ajv: it is a dependency, and this extension has neither dependencies nor a
build step.

test/fixtures/diagrams/corpus.json holds 39 broken specs with the exact outcome
of each attempt. They are seeded from the ways models are known to get this
wrong, not collected in the field; real ones should be added as they turn up.

diagramSchema (18 tests) and diagramRepair (65).
…house style

  diagram/theme.js    the style guide as numbers: title 15, node name 13
                      semibold, second lines and edge labels 11.5, nothing
                      under 10.5; radius 8, border 1.25, padding 12; groups
                      inset 16; the accent a low-opacity fill and a 2px border.
                      Colours are editor theme tokens, with two fixed palettes
                      for exported files.
  diagram/layout.js   layout(spec, opts) gives geometry; inspect(geometry) lists
                      whatever overlaps, overflows or leaves the frame.

The spec names ELK.js. This is not ELK: ELK is EPL-2.0 in a repository that is
otherwise MIT-clean, about 1.5 MB, and this extension has no build step. The
spec's own open question asks whether a simpler layered layout is needed as a
fallback; this is that layout, as the only one. layout() is the single entry
point, so ELK can replace it without touching anything else.

What it does: ranks by longest path; crossings reduced by sweeps; positions
across the flow solved as constraints and then lined up by medians; a port per
edge; a track per connector in each gap, so lines cross but never run along
each other; and each edge label placed BESIDE its line, scored against nodes,
lines and other labels. A flow asked for left-to-right that will not fit its
column is laid out top-to-bottom instead.

A group's name sits in the top-left of its frame — which, when the flow runs
down, is exactly where lines come in. No connector runs through it. The name
slides along the frame to the nearest clear stretch; if there is none and the
lines have under 72px to move, they are moved to the far side of the name (the
frame grows by that much, never past the column); otherwise the name stays and
is marked to be drawn over the line.

Measured over 6,000 random specs at the model-facing limits: no node overlaps
another, no connector crosses a node or an unbacked group name, nothing leaves
the frame. Three known limits, each bounded in the suite so it cannot quietly
get worse:

  an edge label that falls back to a halo over a line   6.5% of specs
  two connectors that swap lanes in one gap             0.3%
  a group's name with a connector behind it             2.7% of specs with groups

None of the nine gallery diagrams (fixtures/diagrams/gallery.json) shows any of
the three, at its natural width or at 560, 420 and 320px. The twelve-node one
lays out in under a millisecond once warm; the slowest of the 6,000 took 7ms.

diagramLayout (23 tests, one of them a 1,500-spec fuzz).
  diagram/scene.js   geometry to a tree of elements, built from two allow-lists:
                     seven tags (svg g rect path text title style) and a set of
                     attributes with no href, no style, no id and no on*.
                     mount() turns it into DOM with createElementNS and
                     textContent; toSvg() into an escaped string for a file. A
                     label is never parsed as markup by either, and both refuse
                     anything off the lists rather than trust whoever built the
                     tree.
  diagram/text.js    the outline a screen reader is given (nodes, then edges, in
                     reading order); the one-line stub that stands in for a spec
                     once a conversation is compacted; Mermaid and source
                     export.
  diagram/ascii.js   for when a picture cannot be shown: the SAME layout on a
                     character grid, so it puts things where the picture would.
                     Plain ASCII; East Asian wide characters and emoji count as
                     two cells, combining marks as none.

A linked node carries data-lc-link = its NODE ID, never the path: whoever
handles the click looks the link up in its own copy of the spec.

A group's name that the layout could not keep clear of a connector is drawn
after the connectors, with the halo an edge label gets, so the line reads as
passing behind the word.

Snapshots of three gallery diagrams in both palettes (fixtures/diagrams/
snapshots): a change to the style, the layout or the painter shows up there as
a diff. diagramScene (23 tests) and diagramText (17).
Ask the agent how something is put together and it answered with boxes and
arrows made of characters, which break with the font, the theme and the width
of the panel. Now it has a tool. It describes the structure; the editor
validates it, lays it out and paints it in the chat, in the editor's theme,
with nodes that open the code they stand for.

Graph JSON only: phase 1 of the spec. Mermaid, Vega-Lite and raw SVG are not
built, and the rules tell the model to use a table for numbers and a numbered
list for sequences rather than a format the chat would show as source.

The model's side (diagram/tool.js)
  render_diagram   its input schema is the validator's own schema object
  the rules        when a diagram earns its place; never draw with characters;
                   asked for a diagram, draw it here rather than writing a file
                   of diagram source; the title states the takeaway; one
                   accent; split above twelve nodes; link nodes to workspace
                   files. One worked example, itself a valid spec.
  the answer       {"ok":true,"id":"d-1"}, or every error at once
  About 520 tokens of tool and 450 of rules on every request from a client
  that can render — constant for a session, so they sit in the cached prefix.

One conversation's diagrams (diagram/service.js)
  A call that fails validation is sent back ONCE, with every error. A second
  failure is drawn degraded, with a banner naming each loss and a Retry button
  that the user presses, not the editor. Three bounces in one run and every
  later call is final. A diagram still owed its repair when the run ends — the
  model gave up, ran out of steps, was stopped — is settled from its last
  attempt before agentDone, so no placeholder is left waiting. Arguments cut
  off at the token limit are asked for again, never repaired.

In the chat (media/chat.html)
  A placeholder from the moment the call starts, titled as soon as the title
  has streamed; then the picture. The shared modules are inlined into the page
  by diagram/bundle.js under the page's existing nonce: the content security
  policy is unchanged, and the page still loads nothing. Text is measured in
  the real font. A column too narrow for a left-to-right flow turns it
  downward, and it turns back when the column widens. A full-size view with
  zoom and pan. The "auto-fixed" badge and its details; the degraded banner;
  the errors and the source when nothing can be drawn. Copy source, SVG, PNG,
  open as Mermaid, insert into a Markdown file. The title is the accessible
  name, and a text outline is there for screen readers. If painting fails, or
  the modules never loaded, the same diagram is shown as text — drawn by the
  page in the first case and by the host in the second.

What the page may ask of the editor
  openLink, export, retry and ascii. Each names a diagram by id and is looked
  up in the host's own records. A click sends the NODE id; the path comes from
  the record and is resolved again at that moment (diagram/links.js): inside a
  workspace folder once ".." and symlinks are followed, and a file. SVG and PNG
  come back from the page, so they are checked before they are written
  (diagram/exportCheck.js). The spec lists two actions; its own Retry button
  and text fallback need the other two.

Sessions and context
  The final record — spec, status, fixes, what was lost — is stored with the
  turn and replayed when a chat is reopened; nothing is repaired again. A
  session file is input too: a stored spec passes the validator on its way
  back in, in the host and in the page. At compaction a diagram becomes one
  line ("diagram: <title>, 4 nodes, id d-17"), and get_diagram is offered from
  then on. A session export writes each diagram as a fenced Mermaid block.

Who gets it
  A run carries client.render: rich or ascii. Rich is the chat webview, unless
  levelcode.ai.diagrams.enabled is off or the model's catalog row says
  diagrams: false — the switch for a model that fails the eval. An ascii client
  is offered neither the tool nor the rules. Chat mode has no tools and is
  unchanged.

Counters (diagram/stats.js, "AI: Diagram Statistics")
  First-pass valid rate, fixes by rung, error classes per model, render time,
  tokens per diagram, and answers that drew with characters anyway. Kept in
  the editor's own storage and sent nowhere; enums and numbers only, never a
  label, a title or a path.

Providers: onToolStart now carries the call's id, a new onToolInput streams its
arguments, and a turn returns the raw text of arguments that did not parse.
All additive.

Four existing suites change. agentNoWorkspace and sessionsUi pinned the shape
of the tool list in source, and now pin the new shape. authRetryCallers and
sessionExpiredHost slice compactAgentMemory and resumeSession out of
extension.js, and needed the names those now use.

Tests: diagramAgent (22 — the real runAgent with a scripted provider),
diagramSession (12), diagramLinks (9), diagramHost (27 — extension.js's own
functions against a stand-in for vscode), diagramStats (8), diagramUi (11).
…, and an eval for models

Three scripts for the three things the suites cannot show. None is in the
gate, which is plain Node and stays that way.

scripts/diagram-browser-check.js — the shipped chat.html in headless Chrome,
built the way the host builds it (the same bundle, the same policy) and fed
records made by the real service. 155 checks across both themes: every diagram
is an SVG of the painter's elements only; hostile labels are on the page as
text and nothing ran; no policy violation and no request; the placeholder, the
badge, the banner, the failed view; a click names a node, never a path; SVG
and PNG export pass the host's own check; and a painter that throws, or a page
with no diagram modules at all, both end in the text fallback. Each step waits
for its result, not for a length of time: the first version waited, and failed
one run in three.

scripts/diagram-editor-check.js — the real editor. A second, throwaway instance
of the dev build (its own profile, sessions folder and workspace) loads THIS
checkout's extension in place of the built-in one, talks to a stand-in provider
on localhost, and is driven through the DevTools protocol. 17 checks: the tool
and its rules reach the model; the diagram is painted in the real webview in
the editor's colours; a wider column re-lays it out; a linked node opens its
file at the symbol; the session on disk holds the diagram; and the answer
around it, which quotes a code fence, is one paragraph (the chat fix earlier in
this branch).

It exists because of a mistake worth writing down. The feature was first called
done having only ever run in a browser, with instructions for trying it that
could not have worked: the work sat uncommitted in a git worktree, and a
worktree has no vscode/ — run-dev.sh runs the extensions of the checkout that
does. The editor that was then tried was develop's, with no diagrams in it.
What does work, by hand, from the checkout that has vscode/:

  ./scripts/run-dev.sh --extensionDevelopmentPath=<checkout>/extensions/levelcode-ai

scripts/diagram-eval.js — the spec's eval: 30 prompts that should produce a
diagram and 10 that should not (fixtures/diagrams/eval-prompts.json), each run
through the real agent loop and scored on format choice, first-pass validity,
fixes by rung, degraded rate, error classes, tokens, node-count overruns, title
quality and answers that drew with characters. --dry-run uses a scripted model
and no network. --run makes billed calls on your key, says how many before the
first one, and is refused without a model and a key; the workspace is an empty
temporary folder and every approval is answered no.

It has not been run against any model. That is the next thing this feature
needs: nothing here says how often a real model's first attempt is valid, or
whether it reaches for the tool at all.

diagramEval (16 tests) holds the harness to its own arithmetic: a scripted
model whose mistakes are known comes out at exactly the numbers its script
implies, and nothing is sent without --run, a model and a key.
…ow to check it

docs/RICH-DIAGRAMS.md is the implementation record for the Rich Diagrams spec
(2026-10-04), under the spec's own section names, because the code cites them.
Per section: the rule, what the code does about it, and each place the build
chose differently, with the reason, so the choice can be argued again. A status
per requirement; the known limits with their measured rates; what is verified
and what is not — no live model has drawn a diagram yet, and the eval has not
been run.

It is not the spec. The spec is a private document; its links, and one line
about pricing, are left out of a public repository.

CLAUDE.md gains the feature's entry and three conventions that cost time to
learn: the shared diagram modules are pasted into a script block and must never
spell a script tag or an HTML comment opener; the host suites' brace matcher
cannot read a backtick inside a regex literal; and work in a worktree is not in
the editor until it is loaded with --extensionDevelopmentPath.

test/fixtures/diagrams/README.md says what each fixture is for, and how to add
a real broken spec to the corpus.
Copilot AI balanced review requested due to automatic review settings October 5, 2026 04:38
Comment thread extensions/levelcode-ai/diagram/layout.js Fixed
Comment thread extensions/levelcode-ai/diagram/theme.js Fixed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Untrusted schema versions can crash validation, and several repair, replay, compaction, statistics, and accessibility paths remain incorrect.

Review effort: Balanced
Findings: 2 High severity · 5 Medium severity

Open (7)
What changed in this PR

Adds phase-one rich Graph JSON diagrams to agent chat, including validation, rendering, persistence, export, accessibility fallbacks, and evaluation tooling. It also fixes streamed Markdown fence handling.

Changes:

  • Adds the render_diagram agent tool and provider streaming support.
  • Implements safe diagram layout, rendering, links, exports, persistence, and statistics.
  • Adds extensive unit, browser, editor, snapshot, and evaluation coverage.
File Description
CLAUDE.md Documents diagram architecture and conventions.
docs/​RICH-DIAGRAMS.md Records the feature design.
extensions/​levelcode-ai/​agent.js Integrates diagram tools into agent runs.
extensions/​levelcode-ai/​diagram/​ascii.js Adds text fallback rendering.
extensions/​levelcode-ai/​diagram/​bundle.js Bundles shared modules into the webview.
extensions/​levelcode-ai/​diagram/​exportCheck.js Validates exported SVG and PNG data.
extensions/​levelcode-ai/​diagram/​layout.js Implements layered diagram layout.
extensions/​levelcode-ai/​diagram/​links.js Resolves safe workspace links.
extensions/​levelcode-ai/​diagram/​repair.js Implements normalization and repair.
extensions/​levelcode-ai/​diagram/​scene.js Builds allow-listed SVG scenes.
extensions/​levelcode-ai/​diagram/​schema.js Defines the versioned Graph JSON schema.
extensions/​levelcode-ai/​diagram/​service.js Manages diagram lifecycle and records.
extensions/​levelcode-ai/​diagram/​stats.js Collects local rollout metrics.
extensions/​levelcode-ai/​diagram/​text.js Generates outlines, stubs, and Mermaid.
extensions/​levelcode-ai/​diagram/​theme.js Defines diagram styling and theme tokens.
extensions/​levelcode-ai/​diagram/​tool.js Defines tools and model instructions.
extensions/​levelcode-ai/​diagram/​validate.js Validates schema and semantics.
extensions/​levelcode-ai/​extension.js Wires host actions, persistence, and exports.
extensions/​levelcode-ai/​media/​chat.html Adds diagram cards and zoom UI.
extensions/​levelcode-ai/​package.json Registers settings and statistics command.
extensions/​levelcode-ai/​providers/​anthropic.js Streams tool IDs and partial arguments.
extensions/​levelcode-ai/​providers/​catalog.js Adds per-model diagram capability.
extensions/​levelcode-ai/​providers/​index.js Forwards diagram streaming callbacks.
extensions/​levelcode-ai/​providers/​openaiCompat.js Handles streamed OpenAI tool arguments.
extensions/​levelcode-ai/​providers/​translate.js Preserves malformed raw tool arguments.
extensions/​levelcode-ai/​scripts/​diagram-browser-check.js Adds browser-level rendering checks.
extensions/​levelcode-ai/​scripts/​diagram-editor-check.js Adds real-editor integration checks.
extensions/​levelcode-ai/​scripts/​diagram-eval.js Adds model evaluation harness.
extensions/​levelcode-ai/​sessionEvents.js Stores and exports diagram events.
extensions/​levelcode-ai/​sessions.js Persists diagrams with sessions.
extensions/​levelcode-ai/​test/​agentNoWorkspace.test.js Updates host-gated tool assertions.
extensions/​levelcode-ai/​test/​authRetryCallers.test.js Stubs diagram compaction behavior.
extensions/​levelcode-ai/​test/​chatMarkdownFences.test.js Tests streamed Markdown fence handling.
extensions/​levelcode-ai/​test/​diagramAgent.test.js Tests agent diagram integration.
extensions/​levelcode-ai/​test/​diagramEval.test.js Tests evaluation logic.
extensions/​levelcode-ai/​test/​diagramHost.test.js Tests host wiring and security.
extensions/​levelcode-ai/​test/​diagramLayout.test.js Tests layout and fuzz cases.
extensions/​levelcode-ai/​test/​diagramLinks.test.js Tests workspace-link containment.
extensions/​levelcode-ai/​test/​diagramRepair.test.js Tests the repair ladder.
extensions/​levelcode-ai/​test/​diagramScene.test.js Tests SVG safety and snapshots.
extensions/​levelcode-ai/​test/​diagramSchema.test.js Tests schema validation.
extensions/​levelcode-ai/​test/​diagramSession.test.js Tests persistence and replay.
extensions/​levelcode-ai/​test/​diagramStats.test.js Tests metric aggregation and privacy.
extensions/​levelcode-ai/​test/​diagramText.test.js Tests textual representations.
extensions/​levelcode-ai/​test/​diagramUi.test.js Tests webview integration and safety.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​README.md Documents diagram fixtures.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​corpus.json Adds malformed-spec corpus.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​eval-prompts.json Adds model evaluation prompts.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​gallery.json Adds representative diagrams.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​arch.dark.svg Adds dark architecture snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​arch.light.svg Adds light architecture snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​decision.dark.svg Adds dark decision snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​decision.light.svg Adds light decision snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​jev.dark.svg Adds dark routing snapshot.
extensions/​levelcode-ai/​test/​fixtures/​diagrams/​snapshots/​jev.light.svg Adds light routing snapshot.
extensions/​levelcode-ai/​test/​sessionExpiredHost.test.js Adds diagram service test stub.
extensions/​levelcode-ai/​test/​sessionsUi.test.js Updates session integration assertions.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread extensions/levelcode-ai/diagram/stats.js
Comment thread extensions/levelcode-ai/diagram/validate.js Outdated
Comment thread extensions/levelcode-ai/agent.js
Comment thread extensions/levelcode-ai/diagram/service.js
Comment thread extensions/levelcode-ai/media/chat.html
Comment thread extensions/levelcode-ai/media/chat.html Outdated
Comment thread extensions/levelcode-ai/sessionEvents.js Outdated
…eant to be; a dead store goes

Two findings from the code-quality review. Both are right.

theme.js, charEm() — "the condition c > 0xffff is always false". The width
estimate used where there is no canvas asked "CJK and other full-width scripts"
(U+2E80 and up, 1.0 em) before "emoji / astral" (above U+FFFF, 1.1 em). Every
code point above U+FFFF is above U+2E80 too, so the second question was never
reached and an emoji was measured at 1.0 em. The estimate is meant to err wide
— the failure that matters is text overflowing its box — so the questions are
now asked the other way round.

Nothing on screen changes. The chat measures with the real font, and no fixture
contains a character beyond the basic plane, so no snapshot moves. What changes
is the fallback when the canvas refuses the font, and the estimates made in
Node.

layout.js — "the value assigned to span here is unused". When several
connectors share a side too short for them, the node grows to hold exactly what
they need; span was recomputed afterwards and never read. The recompute is
gone and span is a const. it.cs, set on the same line, IS read later, and
stays. 6,000 random layouts and the nine gallery diagrams are byte-identical to
the code before.

diagramLayout gains a test for the estimate's classes: narrow, ordinary and
wide letters, digits and capitals, a CJK character at a full em, an emoji at
1.1 em and counted as one character. It fails on the old order, on exactly the
condition the review named.

62 suite files pass on macOS and in a Linux container.
…model, a version or a shape

Two findings from the Copilot review, both right, and a third of the same kind
that looking for them turned up.

stats.js — "IDs such as `constructor` resolve inherited properties". The
counters are plain objects keyed by model id, and a model id is whatever a
provider calls it. For a model called "constructor" the lookup found Object
itself: `calls++` wrote to Object, and the next line threw. For "__proto__" it
found the prototype every object shares, and `calls++` wrote there — every
object in the extension host had a `calls` of NaN — before it threw. The service
swallows a counter that throws, so the diagram was still drawn; the write was
not undone. An error class is a key too: "constructor" passes the class
pattern, and its count came out as "function Object() { [native code] }1".

A key is now read as an own property and written as one (defineProperty, so
"__proto__" is an entry and not a call to the prototype's setter). The store
stays an ordinary object: what is already in the editor's storage reads as
before, and it survives the JSON round trip it is stored through.

validate.js — "a crafted schema version such as `toString` resolves through
SCHEMAS' prototype". validate() promises a list of errors and threw instead:
SCHEMAS["toString"] is a function, so the version was "known", and the check
that followed read `.type` of undefined. prepare() threw with it, and so did
accept(), which is what reads a stored record back. A version is now looked up
only when it is one of the numbers this editor reads. That is deliberately
stricter than "an own key of the registry": no schema declares `v`, so the
string "1" passed as version 1 and was stored as a string, with nothing after
the lookup to object. On the ladder a version written as digits is still read
as the number, as before.

repair.js — the same lookup a third time, found by feeding those names through
every field. The synonym tables are objects, and `"shape": "constructor"`
found Object: the spec came out with a FUNCTION for a shape, and the model was
told "/nodes/0/shape: expected a string, got function" about a string it had
written. A line style of that name was read as solid without a word. A word is
now looked up as the table's own entry, so this one is unknown like any other
unknown word: drawn with the default, and said so.

Nothing else in these modules indexes an object with text from a spec: ids,
labels and groups are kept in Maps, and the schema is walked by its own field
names. The chat page keeps a table of its own, keyed by tool-call id; that one
is fixed with the page, two commits on.

Each has a test that fails on the code as reviewed: six such names as model
ids and one as an error class, with nothing written to Object or its
prototype, before and after a trip through JSON; fifteen things that are not
version 1 — the string "1" and the list [1] among them — as one error each and
never a throw, and four of them up the ladder and through accept(); seven such
words as a direction, a shape and a line style, with exactly the outcome of
"zigzag".
…laced drawing is not handed back, and a replay keeps the order

Three findings from the Copilot review. All three are right.

agent.js:1041 — "this discards parsed.value and passes the raw string … the
original is degraded and a second card is drawn". The service decides whether a
call is the waiting diagram's repair by its title and its node ids, and it read
them off the arguments as they arrived. Arguments that are loose JSON arrive as
text — no title, no nodes — so a repair written with a trailing comma was a
"different diagram": the first attempt was settled as a degraded picture of its
own and the correction was drawn as a second. The same went for a first attempt
in loose JSON (its repair never found it), for a spec inside a `{"spec": …}`
envelope, and for one that says `name` for `title`.

The identity is now read from the spec as the ladder reads it — normalize(),
the lossless rung — whatever form it came in. agent.js is unchanged: it still
hands over the text, so the loose JSON is still counted as the auto-fix it was.
The placeholder's title comes from the same reading, so a call that arrived as
text has one too. And a second drawing of the same title in one run is
recognised as a redraw when it arrives as text.

service.js:260 — "this emits stubs for records that a later drawing replaced".
A redraw replaces the earlier drawing in the chat and a reopened session leaves
it out, but at compaction the model was handed a stub for each — so it could
see, fetch and edit a diagram the user no longer has. A replaced drawing now
gets no stub, wherever its replacement is in the conversation; it is not among
the known ids; and get_diagram, asked for it by an id the model may still
remember, answers with the id of the drawing the chat shows — the end of the
chain when it was redrawn more than once. A record that names itself, or two
that name each other, is a damaged file and not a reason to hang.

sessionEvents.js:149 — "all text blocks are collapsed and emitted before the
diagram". A message that says something, draws, and says more was replayed as
all of its prose and then the picture. It is now walked block by block: prose,
picture, prose, as it was on screen. Every piece after the first carries
`cont: true`, so a reader of the turns can still tell where a message ends
(the chat page does not need it: it already gives a run of answers one
speaker label). A message with no picture in it comes out exactly as it did,
one turn of all its text — and so does every caller that passes no records.

The Markdown export is built on those turns, and getting its order right
showed two more things wrong with it. A picture with no prose before it was
written under whoever spoke last — after a question, under "You". And once a
message is split, the prose after its picture would have been a turn of its
own. The export now works message by message: one rule and one speaker label
per message that says something, its pictures where they stood; a message that
is only a picture joins the answer above it, as before, unless there is none;
a diagram that was never drawn is dropped only after the pieces are grouped,
so the prose after it is not appended to the message before; and the count in
the header is still the number of messages that said something.

Tests, each failing on the code as reviewed. Through the real agent loop: a
repair in loose JSON, a first attempt in loose JSON, three envelope and
renamed-field shapes, a repair with the title reworded, and a redraw in loose
JSON — one diagram on file each time — and a different diagram in loose JSON,
which must NOT be taken for the same one. Three drawings of one diagram leave
one stub, and none in a stretch that holds only a replaced one; the same after
a reload, and in a file whose records name each other. A message of prose,
picture, another tool's call, prose and picture replays in that order with the
same words, and is exported as expected in five shapes.
… is not a button, and one diagram is one card

Two findings from the Copilot review, both right. Checking them in a browser,
with the messages in the order a real run sends them, showed three more faults
in the same code: the page's bookkeeping of which card a diagram is drawn in.

chat.html:1523 — "this is declared modal, but opening it neither makes the
background inert nor traps Tab focus". The full-size view said
aria-modal="true" and covered the page, and the chat behind it could still be
tabbed into. While it is open, every other child of the body is now inert —
not focusable, not clickable, not read out — and released when it closes,
before the focus is handed back (an inert element cannot take it). Tab goes
round inside: past the last stop is the first, before the first is the last,
and a focus that got outside is brought back. A linked node in the view could
be tabbed to and did nothing on Enter; it opens its file now, as a click does.

chat.html:4442 — "the stage is exposed as a button while linked SVG nodes
inside it are independently focusable links". A button's content is no place
for other controls. The picture is no longer a control: no role, no tab stop,
no name. Opening it full size is a button of its own, first in the toolbar
("Full size"), and that is what the keyboard and a screen reader use; a click
on the picture still does the same, as a shortcut for the pointer. Where there
is no picture — the text fallback — there is no such button.

The three it turned up. agent.js announces every render_diagram call the
moment it starts streaming, before anyone can know whether it is a new
diagram, a repair or a redraw, and the page's checks had never been played in
that order.

- A redraw left BOTH drawings on screen, where a reopened session shows one.
  The new call had made a placeholder of its own, the record was drawn in it,
  and the drawing it replaces stayed. The earlier one now goes when its
  replacement arrives, and the redraw stands where the model drew it again —
  whether or not the call announced itself.
- A different diagram arriving while one was being repaired painted over it.
  The page lets the next call take over a card that is waiting for a repair,
  which is right when it is the repair. When it is not, the host settles the
  waiting diagram into that card, and the newcomer then drew over it: two
  diagrams on file, one on screen. A record now has a claim on a card that is
  still a placeholder, or that shows the drawing it replaces, and on no other.
- A tool call named `__proto__` broke the ones after it. The table of cards
  was a plain object keyed by tool-call id, a name a provider chooses.
  `__proto__` made that card the table's prototype; the next lookup of
  `parentNode` threw "Illegal invocation", and that diagram was never shown.
  The table has no prototype now.

Tests. The gate cannot boot the page, and until now it only pinned its source;
every pin was green through all five faults. Two parts of the page are now RUN
there, sliced out of the file. The card bookkeeping, against a stand-in for
the few DOM calls it makes: a repair, two calls in one turn, a redraw with and
without a placeholder of its own, a change of subject, a record with no claim
on the card it points at, seven tool-call ids that are also names the table
could inherit, a wiped transcript, and the sweep at the end of a run — each
judged by what is on screen and where. And the function that decides where Tab
goes.

The browser check gains the modal's behaviour in real Chrome — inert, a focus
that cannot leave, Tab at both ends, Enter on a node, Escape and the focus back
on the button — and a page that plays seven runs in the real order, then
replays each as a reopened session and compares the two. 186 checks, up from
155. The real-editor check opens the view from its button in the real webview
and confirms the page behind it is inert there: 19 checks, up from 17.
…ng only the chat can open

Not from the review: found while fixing what it said about the picture and the
linked nodes inside it.

A linked node is a focusable `role="link"` group carrying `data-lc-link`. In
the chat, a click or Enter asks the host to open the file. "Save as SVG" wrote
the same attributes into the file, where nothing can answer them: opened in a
browser, the picture had tab stops that announced "Open agent.js at runAgent,
link" and did nothing.

A standalone build — the one made for a file — now leaves the four attributes
off. The node keeps what a picture can use: its file icon, and the tooltip
with the path. The picture in the chat and its full-size view are unchanged.

The two arch snapshots lose exactly those attributes on their two linked nodes
and nothing else; the other four snapshots do not move. diagramScene walks the
standalone tree for all four, and checks both nodes are still marked as files.
RICH-DIAGRAMS.md after the #104 review round.

What it now says about the feature: which call counts as a diagram's repair
(its title or its nodes, as the ladder reads them); that a reopened chat shows
what the live one showed — prose and pictures in the order they were written,
one diagram one card, a replaced drawing gone from both; that a replaced
drawing is not handed back to the model; that the full-size view is modal and
is opened by a button of its own; that a saved SVG carries no links; and that
a name arriving from outside is never used to index a plain object by itself.

One new known limit: a repaired diagram is drawn live where its first attempt
was waiting, and a reopened session shows it at the call that got it right.

And the numbers, as run on this tree: 270 tests in the twelve diagram suites
(from 252), 186 checks in the browser (from 155), 19 in the real editor (from
17), the gate on macOS and in the Linux container, the suites with every timer
delayed, and 69 single-edit mutations against this round's changes, all caught
after two gaps were closed. The first mutation set was not run again, and the
doc says so. The count of suite files in the whole gate is gone from the
sentence about it: it goes stale with every branch that adds a test file.
@ndemianc
ndemianc merged commit df63e42 into develop Oct 8, 2026
2 checks passed
@ndemianc
ndemianc deleted the feat/rich-diagrams branch October 8, 2026 03:58
ndemianc added a commit that referenced this pull request Oct 8, 2026
One conflict, in CLAUDE.md: develop (#104) added bullets around the
"Commit ..." line that this branch edits to list `modules/`. Both are
kept — the merged file is develop's plus exactly this branch's changes.

The gate (scripts/test-extensions.sh) passes on the merged tree: 64 test
files, develop's diagram suites and this branch's module suite included.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants